Research

SWAR: Evaluating AI-supported risk-of-bias assessment in systematic reviews

Overview


The use of artificial intelligence in evidence synthesis is developing rapidly, but there remains limited evidence about how accurately, consistently and transparently AI can perform complex research tasks.


This Study Within A Review (SWAR) examines the use of generative AI to support risk-of-bias assessment within a systematic review. It compares assessments produced by human reviewers with those produced by AI using a structured and sequentially parameterised version of the Cochrane Risk of Bias 2 tool.
The project aims to identify where AI-supported assessment may improve efficiency, where discrepancies arise, and what safeguards are required to ensure that researcher judgement, transparency and methodological rigour are maintained.


Methods


The SWAR is embedded within an ongoing systematic review and uses the same included studies and risk-of-bias materials as the main review.
Human reviewers complete risk-of-bias assessments using the standard review process. AI assessments are completed separately using a predefined protocol designed to guide the AI sequentially through the relevant Risk of Bias 2 domains, signalling questions and decision rules.


Assessments are then compared to examine:


agreement between human and AI judgements;differences in responses to individual signalling questions;instances in which the AI applies irrelevant or non-applicable sections;the interpretation and use of supporting evidence;the consistency and reproducibility of AI-generated assessments; andthe circumstances in which human review or correction remains necessary.
The current evaluation focuses on the effect of assignment to intervention and follows the appropriate intention-to-treat estimand within the Risk of Bias 2 framework.


Outputs


Planned outputs from the project include:

  • A structured protocol for using generative AI in Risk of Bias 2 assessment;
  • A sequentially parameterised AI assessment framework;
  • An analysis of agreement and discrepancies between human and AI reviewers;
  • Practical recommendations for transparent and reproducible AI-supported evidence synthesis;academic publications and conference presentations;
  • A living project journal documenting the development, testing and refinement of the approach.


SWAR Real-Time Journal


The project journal provides a developing record of how the AI protocol was designed, tested and revised. It documents methodological decisions, problems encountered, discrepancies between human and AI assessments, and lessons learned during the project.


Follow on with our live journal here


Findings


The project is currently in progress.
Early testing has demonstrated that generative AI can follow structured risk-of-bias assessment processes, but that performance depends heavily on the clarity, order and conditional logic of the instructions provided.


Initial work has also identified several areas requiring careful control, including:


  • Completing sections that are not applicable to the study or assessment;
  • Making judgements without sufficient supporting evidence;treating missing information as evidence of low risk;
  • Failing to distinguish extracted evidence from interpretation;
  • Producing apparently confident conclusions where uncertainty remains.

These findings are informing the continuing refinement of the sequential parameterisation and verification process.

Full findings will be added following completion of the comparative evaluation.